fix(qa): give the expression ledger a fail-policy for "nothing evaluates this" - #16660
Merged
Merged
Conversation
…tes this" `FAIL_POLICIES` had four members and all four describe what an EVALUATOR does when the expression is bad. Five rows classify slots that have no evaluator at all, so each had to borrow a member claiming something stronger: four carried `compile-error` (the Zod parse is the refusal — true, and a property every row in the ledger shares) and `cel-advanced-policy` carried `fail-closed`, a RUNTIME refusal, on a slot whose own `enforcement` cell reads `(no runtime consumer yet)`. On a security-flavoured row that reads as a security guarantee. Adds `unevaluated` and re-states all five rows with it. Every other row's `failPolicy` was re-read against the new member and left where it was. Minting a word that could be borrowed as loosely would reproduce the defect one member wider, so the member arrives with a pin. "Non-empty runtime enforcement" cannot be checked as `enforcement !== ''` — `ExprSurface` makes the cell required, so every row has one. The checkable question is what the cell SAYS: an `unevaluated` row must state the absence it claims and must not name a runtime evaluator site, and it cannot be `state: 'enforced'`. The detector carries a positive control against the COMPILE rows another pin already requires to name the canonical compiler, so an emptied regex reds instead of passing vacuously. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01YFY46JydE1gMxQG1TqBcMZ
Contributor
📓 Docs Drift CheckNothing in this diff resolved to a documentable surface (no symbol, route or SDK anchor derived from 0 changed package(s)), so this run has no opinion about the docs. What this run could not see
Coarse fallback — 0 page(s) merely mention a changed package (the pre-#9192 predicate, kept for the deliberately-wide backstop): |
os-sales
marked this pull request as ready for review
September 7, 2026 16:59
os-sales
enabled auto-merge
September 7, 2026 16:59
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #15533
Ruling: option C, director seat, decision batch #73, 2026-09-07, maintainer verbatim 「同意」, recorded at comment
5565954678.FAIL_POLICIESgains an explicit member meaning "nothing evaluates this slot" —unevaluated— and all fivestate: 'experimental'rows are re-stated with it,cel-advanced-policyincluded. Option B (spreadfail-closed) was refused; option A is what shipped and is what this replaces.What changed
Two files, both inside the private
@objectstack/dogfoodworkspace:packages/qa/dogfood/test/expression-conformance.ledger.ts—FailPolicygainsunevaluatedwith a docblock defining it; the five rows are re-stated; two notes whose prose defended the borrowed value are re-written.packages/qa/dogfood/test/expression-conformance.test.ts—FAIL_POLICIESgains the member, and a new pin refuses the shapes the word could be borrowed in.The vocabulary's locality was re-derived from the delivered diff rather than inherited:
git grep FailPolicyandgit grep EXPRESSION_SURFACEreturn hits in those two files only, andfail-soft-logappears nowhere else in the tree. Nothing published moves. Clause-② staysno.The ruling's census was taken on a pre-#15802 tree — re-derived here
The ruling scopes the work as "re-read 20 rows" (4
compile-error· 7fail-closed· 7fail-soft-log· 2throw). #15500 closed on 2026-09-05 via PR #15802, whose whole purpose was to give every declaration its own ratchet key — which splits previously-collapsed keys into separate rows. The ruling comment is dated 2026-09-07T06:20:23Z — two days later — so it could not carry that arithmetic. (Dates verified against the API rather than inherited: #15500closed_at2026-09-05T10:43:34Z; PR #15802merged_at2026-09-05T10:43:33Z, title "fix(qa): give every expression declaration its own ratchet key".)Re-derived on this branch's merge base by importing the ledger and counting, not by reading the ruling:
failPolicycompile-errorfail-closedfail-soft-logthrowunevaluatedThe re-read covers 23 rows, not 20. The three extra rows are
cel-select-option-visible,cel-inline-grid-cellandcel-field-group-section— named as the split-out rows by the ledger's own header, in the paragraph #15802 added: "The rows split out by that change arecel-select-option-visible,cel-inline-grid-cellandcel-field-group-section." Their fail-policies corroborate it arithmetically: twofail-soft-logand onefail-closed, which is exactly the 7→9 and 7→8 delta above.git log: this checkout is shallow (git rev-parse --is-shallow-repository→true), sogit log -Son those row ids returns the oldest commit inside the shallow window rather than the introducing one, and would have named an unrelated PR.The 23 rows carry 46
coverskeys between them.A second stale number, same cause: the ruling names five
state: 'experimental'rows. There are six today. The sixth iscel-inline-grid-cell, and its verdict is keep — see the table below. The ruling's enumeration of five survives the re-derivation; only the arithmetic around it was stale.The five rows re-stated
cron-knowledge-refreshcompile-errorunevaluatednothing evaluates the resultcron-declared-unwired(7 covers keys)compile-errorunevaluatedNO EVALUATOR FOUNDtemplate-prompt(2 covers keys)compile-errorunevaluatedNO EVALUATOR FOUNDtemplate-title-formatcompile-errorunevaluatedPARSE ONLY in this repo; the one named reader is a build-time lint warning that never looks at the template textcel-advanced-policy(2 covers keys)fail-closedunevaluated(no runtime consumer yet)— a runtime refusal claimed on a slot with no runtime consumer, on a security-flavoured rowcompile-errorclaimed the Zod parse was the refusal. That is true and it is empty: the parse judges the value's shape and never its grammar, and every row in this ledger has it.fail-closedoncel-advanced-policyis the one that cost something past legibility, which is why B was refused.compile-erroris now an unpopulated member. All four rows carrying it were the borrowers. Left in place deliberately — see the notes at the end.Every existing row's re-read verdict — all 23
failPolicyrls-usingfail-closedrls-checkfail-closedsharing-conditionfail-closedcel-validationfail-soft-logcel-hookfail-soft-logcel-formulafail-soft-logcel-field-rulefail-soft-logcel-select-option-visiblefail-soft-logcel-inline-grid-cellfail-soft-logunevaluated. The cell names an evaluator it could not reach (the objectui renderer, outside this checkout) and reports a measured write path on which nothing refuses.unevaluatedasserts an absence; this checkout cannot establish one, and asserting it would be the invented cell this ledger exists to prevent, in the other directioncel-uifail-soft-logsettings-visibilityfail-closedevaluateVisibilityrefuses the save (HTTP 400); a real runtime refusalcel-action-param-option-visiblefail-soft-logcel-bulk-action-visiblefail-closedfallback:false) rather than acting on itcel-field-group-sectionfail-closedcel-uicel-row-crud-visiblefail-closedcel-row-crud-disabledfail-soft-logcel-flowthrowcron-job-schedulethrowtoBoundaryJobSchedulethrows, contained at the call site and loggedcron-knowledge-refreshcompile-error→unevaluatedcron-declared-unwiredcompile-error→unevaluatedtemplate-promptcompile-error→unevaluatedtemplate-title-formatcompile-error→unevaluatedcel-advanced-policyfail-closed→unevaluatedTwo axes were re-read alongside
failPolicyand deliberately left alone.mode: the five re-stated rows keepinterpret. ADR-0058 D6 makesmodea property of what the surface is, not of whether anything runs it, and re-opening it is a different vocabulary question this ruling did not put.state: unchanged on every row;unevaluatedis a fail-policy, not a lifecycle claim.The load-bearing half — the new word cannot be borrowed as loosely
The ruling asks that a row on an
experimentalslot carryingunevaluatedand a non-empty runtimeenforcementbe rejected by the test.enforcement" cannot be spelledenforcement !== ''.ExprSurfacemakes the cell required, so every row has a non-empty one — the five honestunevaluatedrows included, whose cells are long prose. A literal emptiness check would reject exactly the rows the ruling is minting the word for.The checkable question is what the cell says. The pin asserts three things of every
unevaluatedrow:NO EVALUATOR FOUND/PARSE ONLY/no runtime consumer, so the claim is reviewable rather than inferred from silence;NAMES_RUNTIME_EVALUATORis a detector built from the evaluator sites this ledger's own cells already name (celEngine,compileCelToFilter,ExpressionEngine.evaluate,evaluateVisibility,croner, …). Saying the words in (1) is not enough if the same cell names the thing that evaluates it;state: 'enforced'— restricting the pin toexperimentalrows would leaveenforcedas the escape hatch.unevaluatedmeans nothing reads the slot;enforcedmeans the platform enforces it. The contradiction is refused directly.Relationship to ADR-0058 D5. D5's matrix is keyed by (when) × (security-relevance) and every tier in it describes what happens at a site that evaluates something; it has no row for a slot with no site.
unevaluatedis therefore a ledger-local extension, and theFailPolicydocblock says so in those words ("ADR-0058 D5 fail-policy tiers, plus the one state D5 has no tier for"). ⛔ No ADR was edited —docs/adr/**is a governed surface, and the ruling scoped this change to the test file.The pin carries its own anti-vacuity control, in this file's existing idiom (
expect(declarations.length).toBeGreaterThan(0)on the sibling pin): an emptied or mistypedNAMES_RUNTIME_EVALUATORwould make assertion (2) pass over everything. The control asserts the detector fires on everymode: 'compile'row — the rows another pin in the same file already requires to name the canonical compiler — so a broken detector reds instead of passing silently.Demonstrated RED — four legs, each on a distinct assertion
Implementation committed first (
6dc219bebe), then ablated; the legs below were re-run against that exact head. Every leg proves the mutation reached disk (blob hash differs from the HEAD blob, injected/removed anchor counts) before the run is believed, restores under anEXIT/INT/TERMtrap usinggit checkout HEAD --(never the bare form, which restores from a polluted index), and proves byte-identity against the HEAD blob afterwards. Exit codes captured before any pipe.HEAD blobs at that commit: ledger
2dbac23c974541fc897a16739770789bda1cf5ff, test02493cbe009bc4c7660493b00632404b782f31e9.cel-advanced-policy.enforcement→'@objectstack/formula celEngine (interpret) via the advanced-policy runner'(experimental + unevaluated + a cell naming a runtime evaluator)c7aa4cd5≠ HEAD'(no runtime consumer yet) — evaluated by @objectstack/formula celEngine'(keeps the disclaimer, names the evaluator)daae2753≠ HEADcel-advanced-policy→state: 'enforced', stillunevaluatedb3ed4f5d≠ HEADNAMES_RUNTIME_EVALUATOR→ a regex matching nothing088cdb5c≠ HEADrls-using, so assertion (2) is not vacuousgit diff HEADemptyLeg A's failure text, verbatim:
Verification
Heavy runs went through
scripts/pm/os-verify-lock.sh; verdicts are quoted from the lock's ownVERDICTline, never a bare exit code.pnpm --filter @objectstack/dogfood exec vitest run test/expression-conformance.test.ts→VERDICT command-exit 0, 1 file / 5 tests passed (4 before this PR; the new pin is the fifth).pnpm --workspace-concurrency=2 --filter '@objectstack/dogfood^...' build), thentsc --noEmit --listFiles→VERDICT command-exit 0, 0error TS, and--listFilesnames both edited files in the program, so the verdict really covers them.turbo ls --affectedagainst this branch's merge base:@objectstack/dogfoodonly.grep -naPover both files finds none, beyondpnpm check:nul-bytes.Gate union
node scripts/pm/dispatch-gates.mjs --repo objectstack-ai/objectstack --commandsderived 48 runnable commands from this diff's two paths; all 48 were run against the final head6dc219bebe, and the ran-list was reconciled with--ran. Results are quoted from each gate's own output, with the exit code captured before any pipe.48 derived · 48 run · 0 NOT-MEASURED · 0 UNRUN, and every one exited 0.
dispatch-gates --ranconfirms it:exit 3is a prerequisite failure, never a pass:pnpm check:dual-build-cjs-loads— first run exit 3,PREREQUISITE NOT MET ... this gate reads built output, naming 8 packages with nodist/. Re-run after the dependency-closure build: exit 0.pnpm check:type-check-debt— first run exit 3 after 744s, a V8FATAL ERROR: Ineffective mark-compacts near heap limitunder the standard--max-old-space-size=4096. Re-run at 6144 (raised deliberately, for that measured OOM, on a box with 14 GB available): exit 0 in 130s. The gate's own text is explicit that its exit 3 "is NOT a pass and NOT a finding".Slowest four, for whoever schedules this next:
pnpm check:query-options-erasure(259s) ·pnpm check:type-check-debt(130s) ·pnpm check:slot-lookup(117s) ·node scripts/check-comment-mask-corpus.mjs(110s).dispatch-gatesreports this tree as behindorigin/mainin files the derivation reads (.github/workflows/lint.yml,package.json). Diffed: main has added exactly two families since this branch point, and neither reaches this diff.pnpm check:pm-widening-tellsis a--self-test-only step whose own comment records that its input is "a DIFF supplied by its caller, never a file in the tree" — checker-health, not a verdict on any diff.pnpm check:scaffold-emission-policydeclaresPOLICY_SOURCE = 'packages/cli/src/commands/init.ts'and two docs.mdxpages as its inputs; this PR touches none of them. The remaining always-runs tail is CI's to answer.@objectstack/dogfoodsuite was not run locally. It boots real example apps ~130 times and CI shards it three ways; this diff touches one test file and its ledger and no runtime source. The suite's own file was run in full. The rest is declared to CI.Changeset
None —
skip-changeset.@objectstack/dogfoodis"private": true; this diff is confined to it and releases nothing from any package. That is the workflow's own stated criterion for the label ("such a PR releases nothing"). The label is applied on this PR at open time, not left for the Check Changeset job to red first.验收备注 — found, deliberately not filed
compile-erroris now an unpopulated member ofFAIL_POLICIES. All four rows that carried it were borrowers, so it now names no row. Retiring it is a live question — [finding]CronExpressionInputSchema/TemplateExpressionInputSchemafix the dialect only on the bare-string arm — the envelope arm accepts any declared dialect, so a cron-typed slot parses{ dialect: 'cel', source }green #15028 records that it reads stronger than it is — but an unused vocabulary member is dead code, which is not a filing category, and retiring it is a decision this ruling did not make. Noted, not filed.The reverse pin was considered and deliberately NOT written. The symmetric rule — a row whose cell declares no evaluator and names none must be
unevaluated— is green on today's ledger (the three disclaimer phrases match exactly the fiveunevaluatedrows and no other row, measured). It is the leg that would have caughtcel-advanced-policyborrowingfail-closedin the first place. It is not here on purpose: its population is defined by matching three phrases of prose, so a future unwired row that spells its absence any other way passes it silently while the pin reads as "the old words can no longer be borrowed". A partial gate that reads as complete is the exact defect family this ledger catalogues, and adding one to close this card would be a poor trade. The forward direction has no such problem — it governs every row carrying a machine-readable field value. Noted, not filed.cel-inline-grid-cellcarries a runtime fail-policy over a surface its own cell says was NOT MEASURED IN THIS REPO. Row 9 above. Itsfail-soft-logis defensible on the measured half (the write path reads nothing, so nothing refuses), and itsnotealready sets the condition for re-stating it. But the ledger has no word for "an evaluator exists and is out of reach", which is a third dialect of the same legibility question this card answers for "no evaluator at all". Adjacent, not this card, and not a defect against a declared contract. Noted, not filed.A pre-existing editing artifact in the ledger header, untouched. Lines 31-32 of
expression-conformance.ledger.tsread "Two limits survive and / One limit survives and is worth knowing…" — a half-replaced sentence left by an earlier edit, plainly visible in the diff's context lines. It is a documentation nit in a file this PR edits, not a defect in anything the ledger asserts, and this PR's diff is deliberately confined to the vocabulary change. Noted, not filed.Generated by Claude Code